Benchmarking Speech Recognition Models for Medical Consultations in Latin American Spanish: A Comparative Evaluation with Fine-Tuning
This study benchmarks ten speech-to-text models on Latin American Spanish medical consultations, finding that while fine-tuning Whisper Large v3 did not surpass its vanilla version, it outperformed other open-source and closed-source models on key metrics, establishing it as the optimal open-source solution for AI medical scribes in this context.